Papers with end-to-end set

    1 papers
    SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories (2024.emnlp-main)

    Copied to clipboard

    Challenge: Large Language Models (LLMs) have made significant progress in writing code, but can they be used to reproduce results from research repositories?
    Approach: They propose a benchmark to evaluate the capability of Large Language Models to reproduce results from research repositories.
    Outcome: The benchmark aims to capture the realistic challenges faced by researchers working with machine learning and natural language processing repositories.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations